Papers by Khanh Chi Le
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these models reflect human-like cognition versus advanced pattern recognition remains an open question. |
| Approach: | They conduct a series of targeted experiments to assess whether LLMs construct semantic representations and pragmatic inferences in a human-like manner. |
| Outcome: | The proposed framework can be used to assess the cognitive and linguistic capabilities of large language models (LLMs). |
ScholaWrite: A Dataset of End-to-End Scholarly Writing (2026.acl-long)
Copied to clipboard
| Challenge: | SCHOLAWRITE traces the multi-month journey from initial drafts to final manuscripts . authors demonstrate the value of capturing scientists’ cognitive writing process . |
| Approach: | They present a dataset of end-to-end scholarly writing tracing the multi-month journey from initial drafts to final manuscripts. |
| Outcome: | The first dataset of end-to-end scholarly writing traces the multi-month journey from initial drafts to final manuscripts over four months. |
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs (2026.eacl-long)
Copied to clipboard
Karin de Langis, Jong Inn Park, Bin Hu, Khanh Chi Le, Andreas Schramm, Michael C. Mensink, Andrew Elfenbein, Dongyeop Kang
| Challenge: | Working memory is a critical component of human intelligence and executive functioning . it is correlated with performance on various cognitive tasks, including fluid intelligence . |
| Approach: | They apply working memory tasks to large language models to estimate working memory capacity . they find that LLMs exceed normative human scores, but not executive functioning benchmarks . |
| Outcome: | The proposed models do not show higher performance on executive functioning tasks or problem solving benchmarks. |